Papers with moral alignment
The Greatest Good Benchmark: Measuring LLMs’ Alignment with Utilitarian Moral Dilemmas (2024.emnlp-main)
Copied to clipboard
| Challenge: | Our analysis across 15 diverse LLMs reveals consistently encoded moral preferences that diverge from established moral theories and lay population moral standards. |
| Approach: | They propose to evaluate the moral judgments of large language models using utilitarian dilemmas to determine their moral alignment. |
| Outcome: | The findings highlight the ‘artificial moral compass’ of Large Language Models, offering insights into their moral alignment. |
Do VLMs Have a Moral Backbone? A Study on the Fragile Morality of Vision-Language Models (2026.findings-acl)
Copied to clipboard
Zhining Liu, Tianyi Wang, Xiao Lin, Penghao Ouyang, Gaotang Li, Ze Yang, Hui Liu, Sumit Keswani, Vishwa Pardeshi, Huijun Zhao, Wei Fan, Hanghang Tong
| Challenge: | Vision-Language Models (VLMs) have advanced multimodal learning, driving progress in cross-modal reasoning. |
| Approach: | They propose to examine moral robustness of vision-language models by analyzing their moral stances under multimodal perturbations. |
| Outcome: | The proposed model-agnostic multimodal perturbations expose VLMs to a variety of moral vulnerabilities, including a sycophancy trade-off where stronger instruction-following models are more susceptible to persuasion. |